Query Expansion for Khmer Information Retrieval
نویسندگان
چکیده
This paper presents the proposed Query Expansion (QE) techniques based on Khmer specific characteristics to improve the retrieval performance of Khmer Information Retrieval (IR) system. Four types of Khmer specific characteristics: spelling variants, synonyms, derivative words and reduplicative words have been investigated in this research. In order to evaluate the effectiveness and the efficiency of the proposed QE techniques, a prototype of Khmer IR system has been implemented. The system is built on top of the popular open source information retrieval software library Lucene1. The Khmer word segmentation tool (Chea et al., 2007) is also implemented into the system to improve the accuracy of indexing as well as searching. Furthermore, the Google web search engine is also used in the evaluation process. The results show the proposed QE techniques improve the retrieval performance both of the proposed system and the Google web search engine. With the reduplicative word QE technique, an improvement of 17.93% of recall can be achieved to the proposed system.
منابع مشابه
QEA: A New Systematic and Comprehensive Classification of Query Expansion Approaches
A major problem in information retrieval is the difficulty to define the information needs of user and on the other hand, when user offers your query there is a vast amount of information to retrieval. Different methods , therefore, have been suggested for query expansion which concerned with reconfiguring of query by increasing efficiency and improving the criterion accuracy in the information...
متن کاملImproved Skips for Faster Postings List Intersection
Information retrieval can be achieved through computerized processes by generating a list of relevant responses to a query. The document processor, matching function and query analyzer are the main components of an information retrieval system. Document retrieval system is fundamentally based on: Boolean, vector-space, probabilistic, and language models. In this paper, a new methodology for mat...
متن کاملImproved Skips for Faster Postings List Intersection
Information retrieval can be achieved through computerized processes by generating a list of relevant responses to a query. The document processor, matching function and query analyzer are the main components of an information retrieval system. Document retrieval system is fundamentally based on: Boolean, vector-space, probabilistic, and language models. In this paper, a new methodology for mat...
متن کاملPrototyping a Vibrato-Aware Query-By-Humming (QBH) Music Information Retrieval System for Mobile Communication Devices: Case of Chromatic Harmonica
Background and Aim: The current research aims at prototyping query-by-humming music information retrieval systems for smart phones. Methods: This multi-method research follows simulation technique from mixed models of the operations research methodology, and the documentary research method, simultaneously. Two chromatic harmonica albums comprised the research population. To achieve the purpose ...
متن کاملQuery expansion based on relevance feedback and latent semantic analysis
Web search engines are one of the most popular tools on the Internet which are widely-used by expert and novice users. Constructing an adequate query which represents the best specification of users’ information need to the search engine is an important concern of web users. Query expansion is a way to reduce this concern and increase user satisfaction. In this paper, a new method of query expa...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2010